Frontiers in Genetics
○ Frontiers Media SA
Preprints posted in the last 7 days, ranked by how well they match Frontiers in Genetics's content profile, based on 230 papers previously published here. The average preprint has a 0.18% match score for this journal, so anything above that is already an above-average fit.
Zhu, J.; Baousi, A.; Morris, A. P.; Guo, H.
Show abstract
Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study (GWAS) summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs directly from individual-level data, with the aim of improving predictive performance over standard PRSs through their ability to model non-additive genetic effects. However, their superiority across studies has been inconsistent, and the conditions under which they provide meaningful improvements remain unclear. We combined theoretical analysis, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretical analysis showed that standard PRSs can implicitly capture part of the genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects, thereby losing less information than commonly assumed. Although nonlinear models have a higher theoretical potential, their greater flexibility incurs a bias-variance trade-off that can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed the standard PRS only when the genetic architecture involves a sufficiently large proportion of interaction genetic variance concentrated across relatively few interaction effects and large training samples were available. Random forest consistently underperformed the standard PRS. In an application to ischemic heart disease prediction using UK Biobank data, XGBoost showed no meaningful improvement in predictive performance over the standard PRS, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.
Mansoor, R.; Minhas, A. S.; Thomas, A.; Mansoor, A. A.; McCambridge, A. H.; Dilts, C.; Eshak, J.; Govani, D.; Nylin, B.; Trinidad, J. C.; Kanaan, A. Y.; Kara, E.; Fielder, A.; Fielder, I.; Iglendza, A.; Mukatash, Y.; Pumnea, B.; Menzel, M. M.; Shabazz-Henry, A. L.; Niepielko, M. G.; Gao, M.
Show abstract
The QxxR motif is evolutionarily conserved within DEAD-box RNA helicases, including Drosophila Me31B and human DDX6, which post-transcriptionally regulate gene expression during animal development. A pathogenic H372R substitution (QxHR to QxRR) in the QxxR motif of human DDX6 has been associated with various developmental defects, but how this motif contributes to DDX6-family protein function remains unclear. Here, we used Drosophila Me31B as an in vivo model to investigate the QxxR motifs developmental role. We generated a Drosophila strain carrying the corresponding H333R missense mutation in Me31B and characterized its effects on female fertility, embryonic viability, germline development, and Me31B-associated molecular pathways. The me31BH333R mutation reduced female fertility in a gene dose-dependent manner, with homozygous mutant females being sterile. Embryos from the mutant females also exhibited primordial germ cell defects. Despite these developmental phenotypes, the me31BH333R mutation did not significantly alter Me31B protein abundance, global ovarian transcriptome or proteome profiles, or representative germ plasm mRNA and protein localization. In contrast, bait-normalized IP-MS analysis revealed altered enrichment of selected Me31B-associated proteins, including increased association of known Me31B interactors Trailer hitch (Tral) and Ypsilon Schachtel (Yps). These findings establish Me31BH333R as an in vivo model for investigating the conserved QxxR motif and suggest that disruption of this motif compromises development not through broad changes in gene expression, but potentially through altered composition or regulation of Me31B-containing ribonucleoprotein complexes.
Yelgi, A.; Tavangari, S.; Shakarami, Z.; Janfaza, S.
Show abstract
Accurate epigenetic age prediction from DNA methylation profiles is intrinsically high-dimensional, creating a need for parsimonious models that preserve predictive performance while reducing the number of assayed cytosine-phosphate-guanine (CpG) loci. This study introduces MOSurvivor, a population-based multi-objective search framework that jointly optimizes a weight-threshold CpG selector and eight XGBoost hyperparameters. Experiments used the GSE40279 whole-blood cohort (656 individuals profiled on the Illumina HumanMethylation450 platform). After retaining 1,000 age-correlated CpGs, five strategies were evaluated on the same 30 seeded 80:20 train/test splits: fixed-parameter XGBoost using all 1,000 CpGs, random search, a genetic algorithm, particle swarm optimization, and MOSurvivor. Internal fitness was estimated using three-fold cross-validation on each training set. Across the 30 held-out test sets, MOSurvivor achieved a mean absolute error (MAE) of 4.149 {+/-} 0.300 years, root mean squared error of 5.545 {+/-} 0.392 years, and R2 of 0.855{+/-} 0.027 while retaining 211.6 {+/-} 54.8 CpGs. Relative to full-feature XGBoost (MAE 4.095 {+/-} 0.285 years), MOSurvivor reduced the feature set by 78.8% at an MAE increase of only 0.054 years (1.3%). Paired Wilcoxon tests found no significant accuracy difference between MOSurvivor and any comparator (all unadjusted p > 0.05; all Holm-adjusted p [≥] 0.476). The most recurrent locus, cg16867657, appeared in 29 runs, whereas mean pairwise Jaccard similarity was 0.124, indicating a small stable core embedded in multiple near-equivalent feature subsets. MOSurvivor thus offers a competitive accuracy-parsimony trade-off rather than superior absolute accuracy. External validation and leakage-free nested feature preselection remain necessary before biological or clinical translation. Keywords: epigenetic clock, DNA methylation, feature selection, multi-objective optimization, XGBoost, metaheuristics, biological aging.
Hussain, T.; Anothai, J.; Nualsri, C.; Ali, A.; Khomphet, T.
Show abstract
Drought stress is the major yield limiting factor in upland rice production where the moisture availability is highly variable. Understanding and evaluating how upland rice responds to drought stress is critical to improving resilience and yield stability. In this study performance of sixteen upland rice varieties were evaluated under non-stressed, moderately stressed and highly stressed conditions. Drought stress was introduced by irrigating upland rice at 70% and 50% field capacity (FC) whereas non-stress treatment was irrigated at 100% FC. Irrigation in moderately stressed and highly stressed conditions was also withheld for six days at lateral crop stages to observe temporary wilting by inducing a stress interval. Data on agronomic traits of upland rice was collected in three experimental replications. Results indicated that performance of upland rice varieties was significantly altered under stress conditions and highest performance was observed under non-stressed conditions. Yield losses for short duration and long duration varieties ranged 35-60% and 24-62% under moderate stress whereas it ranged 43-78% and 52-73% under highly stressed conditions, respectively. Overall varieties Dawk Kha, Khao/ Sai and Dawk Pa-yawm, indicated higher stability under stressed conditions therefore, these long duration varieties could be used for obtaining better yields under diverse agroclimatic conditions and under unpredicted weather patterns. Short duration Ma-led-nai-fai and long duration Goo Meung Lung and Bow Leb Nahag could be used for acquiring traits for higher tillering and panicle bearing capacity. Short heighted varieties such as Jao Daeng, Sahm Deuan and Ma-led-nai-fai could be used in breeding for short heighted new varieties to overcome lodging concerns. Strong significant association of GMP, STI, MPRO, MHAR with grain yield under non-stressed, moderately stressed and highly stressed conditions indicated that these indices were appropriate for their use as selection criteria for drought resilience.
Montanez-Valverde, R. A.; Kim, V.; Duran-Luciano, P.; Yuan, Y.; Sofer, T.; Kaplan, R. C.; Gallo, L. C.; Talavera, G. A.; Perreira, K. M.; Daviglus, M. L.; Rosas, S. E.; Llabre, M. M.; Elfassy, T.; Li, X.; Isasi, C. R.; Rodriguez, C. J.
Show abstract
Background. The imprecision of current metrics to capture the complex genetic admixture and racial identity among Hispanic/Latino individuals in the United States [US] is a concern. We examined the relationship of self-reported race and genetic ancestry with hypertension [HTN] among Hispanics/Latinos. Methods. Cross-sectional study of the Hispanic Community Health Study/Study of Latinos (HCHS/SOL), including 10,586 Hispanic/Latino unrelated adults. Genetic ancestry: West African [AA], Amerindian [AI], and European [EA]. Self-reported race: White, Black, Native American, or Multiple/Missing (More than one race or Unknown/Not reported/Refused). HTN: systolic (SBP) [≥]130 mmHg, diastolic blood pressure (DBP) [≥]80 mmHg, and/or use of HTN medications. Age- and sex adjusted models were used. Results. Self-reported race was White (38{middle dot}6%), Black (3{middle dot}6%), Native American (4{middle dot}1%), and Multiple/Missing (53{middle dot}7%), with Unknown/Not reported/Refused representing 32{middle dot}7%. Black and White Hispanics/Latinos had the greatest AA (55{middle dot}7%) and EA (69{middle dot}3%) ancestries, respectively. Each 10% AA increase was associated with OR 1{middle dot}15, SBP beta +0{middle dot}9 mmHg, and DBP beta +0{middle dot}7 mmHg. Conversely, each 10% AI increase was associated with OR 0{middle dot}83, SBP beta -0{middle dot}4 mmHg, and DBP beta -0{middle dot}6 mmHg. HTN prevalence was highest among those with Black race or in the highest AA quantile (45{middle dot}6% and 48{middle dot}0%, respectively), and lowest among those with Native American race or in the highest AI quantile (37{middle dot}6% and 26{middle dot}7%, respectively). Conclusion. One-third of Hispanics/Latinos did not self-report race. Black or White self-reporting race did somewhat relate to AA or EA ancestry, respectively. HTN profiles were related to self-reported race and genetic ancestry in this admixed population.
Richards, C.; La Salle, D. T.; Vila Dieguez, O.; Ward, S. R.
Show abstract
Background: Newly available arm angle data offers a new dimension to understand rising rates of arm injury in MLB pitchers. Purpose: To evaluate the relationship between arm angle, pitch characteristics, and elbow and forearm injury in MLB pitchers.<br><br> Study Design: Retrospective cohort study; Level of evidence, 3 Methods: Statcast data from 2020 to 2025 and MLB injured list (IL) data were used to evaluate arm angle and pitch characteristics in relation to elbow and forearm injuries. Results are presented with and without requirements on prior season workload and for same-season and next-season injury incidence. A generalized additive model (GAM) was used to capture non-linear dependence and interactions between selected features and injury incidence to the elbow or forearm. Average marginal effect (AME) odds ratios are reported for main effect terms. Results: N = 3,812 pitcher-seasons were included. 29% pitchers who underwent UCLR did so in the same season as a forearm injury (tmean=44, tmedian=27 days to surgery). Arm angle, fastball usage, and their interaction were the three most predictive features. Arm angle was positively related to incidence of injury (ORmeanAME=1.014), fastball usage was inversely related to incidence of injury (ORmeanAME=0.243), and arm angle moderated the effect fastball usage at high arm angle, where increased usage was no longer protective. Slider velocity (ORmeanAME=1.072), spin rate (ORmeanAME=1.001), and usage (ORmeanAME=2.039) also significantly predicted injury risk. Fastball velocity was not significant in any fit, with ORmeanAME=0.999 across all fits. Fit-level Nagelkerke R2 values ranged from .019 to .052. Conclusion: Fastball usage and arm angle, not velocity, predicted elbow and forearm injury risk among MLB pitchers, and arm angle was the single most predictive feature. The heterogeneity of risk factors as a function of arm angle, and the novelty of MLB arm angle data, may explain why fastball usage has been previously underexplored as a risk factor. Keywords: baseball; arm angle; fastball velocity; fastball usage; spin rate; UCL; ulnar collateral ligament; elbow injury; forearm injury; Statcast
Kouam, C.; Mingle, J.; Alvarez Jerez, P.; Evans, A.; Moller, A.; Baker, B.; Weller, C.; Paquette, K.; Brooks, J.; Grant, S. M.; Ayuketah, A.; Meredith, M.; Palade, J.; Malik, L.; Hise, K.; Raphael Gibbs, J.; Anderson, J.; Ding, J.; Harbert, R.; Fu, Y.; Zheng, X.; Garcia-Ruiz, S.; Gustavsson, E. K.; Blauwendraat, C.; Ryten, M.; Sedlazeck, F.; Ferrucci, L.; Reed, X.; Nalls, M. A.; Cookson, M. R.; Van Keuren-Jensen, K.; Hutchins, E.; Jain, M.; Billingsley, K. J.
Show abstract
Isoform-resolved transcriptomics is fundamental to decoding the molecular complexity of the human brain, yet population-scale long-read RNA sequencing has remained inaccessible due to labor-intensive library preparation, sensitivity to RNA degradation in postmortem tissue, and the absence of integrated, reproducible analysis pipelines. Here we present SALRR (Scalable Analysis of Long-Read RNA-seq), an integrated wet-lab and computational platform designed to overcome these barriers. Automated ONT long-read cDNA library preparation on the Hamilton Microlab NGS STAR platform reduces hands-on time by 67% and enables 24 libraries per operator per day while maintaining performance across RNA integrity values. A modular, Snakemake-based pipeline performs end-to-end processing from ONT signal data to isoform-level quantification, incorporating SIRV spike-in calibration, multi-stage quality control, and stringent isoform validation. Applied to 10 postmortem frontal cortex samples from the North American Brain Expression Consortium, SALRR identified 31,607 high-confidence isoforms from 10,075 genes, including 8,532 novel splice variants absent from GENCODE v49, and complex splicing events systematically missed by short-read sequencing at neurodegeneration-relevant loci, including GBA1, CCNF, CHCHD10, and TREM2. All protocols and code are openly available, providing a scalable, community-ready framework for isoform-resolved transcriptomics in neurodegeneration, aging, and complex brain disease.
Quigley, H.; Gardiner, B.; McDaid, L.; O'Donnell, C.
Show abstract
Autism Spectrum Disorder (ASD) is a heterogeneous neurodevelopmental condition defined by differences in social communication and restricted, repetitive behaviours. As diagnostic criteria have broadened, ASD is now recognised across a wider range of individuals, raising key questions about its structure: does ASD have discrete sub-types, or is it better conceptualised as a continuous, possibly multidimensional, condition? We aim to explore whether a multidimensional continuum model more accurately captures the variability within ASD. We analysed a large SPARK phenotypic dataset of medical history and diagnostic surveys (background history, SCQ, RBS-R; n=36,710 individuals). We apply and compare two traditional statistical approaches, Factor Analysis and Gaussian Mixture Models, with a modern machine learning technique, the Variational Autoencoder (VAE). VAEs reconstructed unseen test data with ~4-fold better accuracy than Factor Analysis, and ~8-fold better accuracy than Gaussian Mixture Models. We identified four stable latent factors across 100 independently trained VAEs. These four dimensions provide an individual behavioural profile that can be visualized using radar-plots, offering a compact way to compare profiles at the person level. Through further analysis, we found evidence for 3 overlapping clusters or subtypes of ASD identified within the 4D latent space. This work aims to inform new ways of modelling ASD using a VAE that will be able to discern between a continuum or a clustered output and that go beyond binary diagnosis, instead reflecting the complex range of trait profiles, with implications for personalised diagnosis and intervention.
Richards, C.; La Salle, D. T.; Vila Dieguez, O.; Ward, S. R.
Show abstract
Background: Sweeping sliders with large horizontal break are hypothesized to be associated with arm injury, and prior work suggests a link between slider usage, arm angle, and injury. Purpose: To evaluate whether arm angle and slider horizontal movement are associated with elbow or forearm injury risk in MLB pitchers. Study Design: Retrospective cohort study; Level of evidence, 3 Methods: Statcast data from 2020-2025 and injury data were used to study the relationship between sliders and arm injuries. Pitchers with at least 30 innings pitched (IP) were evaluated for same- and next-season injury incidence to the Elbow, Forearm, or Elbow/Forearm. A generalized additive model (GAM) related pitch-level variables to injury incidence, and average marginal effect (AME) odds ratios are reported for main effects. Results: All fits were significant at p<.01. Fits for same-season (N=1,957, p=.00008, R2=.066) and next-season (N=2,511, p=.00005, R2=.053) Elbow/Forearm injury were significant at p<.001. Same-season (R2=.061) and next-season (R2=.055) Forearm injury fits had p=.001. Main effects for slider glove-side movement and fastball usage, their interactions with arm angle, and arm angle main effects were the most predictive features. Across the three fits where the slider glove-side movement main effect was significant, odds ratios ranged from 1.03 to 1.07, meaning that each additional inch of slider glove-side movement was associated with a 3% to 7% increase in the observed injury incidence. Furthermore, this effect was magnified at high arm angle and muted at low arm angle. Conclusion: We observed that slider horizontal movement, arm angle, and their interaction were significantly associated with incidence of Forearm and combined Elbow/Forearm injury. This effect was weak or non-existent at low arm angles and strong at high arm angles, suggesting that arm angle moderates the risk of slider horizontal movement, and providing evidence that sliders with large horizontal break (e.g. sweepers) may pose an injury risk. Keywords: baseball; sweeper; slider; arm angle; horizontal movement; elbow injury; forearm injury
Sawyer, G.; Farooq, B.; Birnie, K.; Fraser, A.; Lawlor, D. A.; Sharp, G. C.; Howe, L. D.
Show abstract
Background: Inequalities exist for many health outcomes, but there is limited evidence regarding menstrual symptoms despite their importance for health and wellbeing. We aimed to investigate inequalities in menstrual symptoms according to socioeconomic position and childhood adversity. Methods: In two generations (G0 mothers and G1 offspring) from the Avon Longitudinal Study of Parents and Children (ALSPAC), a UK prospective cohort study, we examined associations of multiple indicators of socioeconomic position (SEP) and adverse childhood experiences (ACEs) with menstrual symptoms (pain, abnormal uterine bleeding, and premenstrual syndrome (PMS) measured 3-8-years post-birth in G0 and 17-21-years-old in G1), using multivariable logistic regression. Samples ranged from 4,828 to 9,335 G0 participants and 1,288 to 2,757 G1 participants depending on the exposure-outcome association. Missing data were addressed using multiple imputation and inverse probability weighting. Results: Financial difficulties were associated with greater odds of menstrual pain (G1 OR 1.41; 95% CI 1.07, 1.86: G0 OR 1.55; 95% CI 1.36, 1.76) and irregular cycles (G1 OR 1.60; 95% CI 1.12, 2.29: G0 OR 1.48; 95% CI 1.27, 1.72) in both generations, as well as with short/long cycle lengths in G0 only. Lower education and manual social class were also associated with these three menstrual symptoms in at least one generation. Conversely, higher SEP was associated with PMS in both generations. Higher cumulative ACEs were consistently associated with menstrual pain (4+ compared to none: G1 OR 2.15; 95% CI 1.48, 3.11: G0 OR 1.52; 95% CI 1.29, 1.80) and irregular cycles (G1 OR 1.92; 95% CI 1.20, 3.09: G0 OR 1.54; 95% CI 1.26, 1.87) but not cycle length. Lower parental education, financial difficulties, and cumulative ACEs were associated with heavy bleeding in G1 offspring only, whereas financial difficulties, own manual social class, and cumulative ACEs were associated with prolonged bleeding in G0 mothers only. Higher cumulative ACEs were also associated with PMS in G1 offspring only. Conclusions: We found evidence of inequalities according to socioeconomic disadvantage and childhood adversity for multiple menstrual symptoms, although some associations were only observed in one generation. Findings suggest that menstrual symptoms are disproportionately experienced by socially and socioeconomically disadvantaged women.
SULAIMAN, M. A.; Oyeyemi, B. F.
Show abstract
Sub-Saharan African populations carry pharmacogenomic alleles poorly represented in the European-derived reference panels underlying most clinical genotyping tools. We present a curated, machine-readable catalog of nine actionable alleles across six pharmacogenes (CYP2D6, CYP2B6, CYP2C9, CYP2C19, CYP3A5, NAT2) with African-specific frequency ranges, functional annotations, and evidence levels derived from reanalysis of 661 high-coverage whole-genome sequences across seven 1000 Genomes Project African populations. Direct comparison against PharmCAT v3.4.0 shows that CYP2D6 produces zero diplotype calls (0/661 samples callable) due to monomorphic reference positions absent from standard variant-only VCF output, a known limitation whose consequences for African allele carriers had not been reported. afripharmagen's reduced-position strategy identifies 243 CYP2D617 and 134 CYP2D629 carriers from the same input. For CYP2B6, CYP2C9, CYP2C19, and NAT2, both tools show concordance of 95-100%. Frequency gradients (CYP2B66: 30-50%; CYP2D617: 15-35% in West Africa; CYP3A5*1: 60-95%) translate directly into prescribing risk for efavirenz, tramadol, tacrolimus, and isoniazid. Pharmacogenomic decision support in African settings must incorporate population-specific allele definitions and input-format-aware strategies.
Layman, C. E.; Morrow, D.; Wheeler, K.; Caron, T. J.; Davis, B. A.; Bergstrom, P.; Vigh-Conrad, K.; Anderson, T. J.; McElfresh, G. W.; Sterner, K. N.; Sadoughi, B.; Snyder-Mackler, N.; Hansen, S. G.; Bimber, B. N.; Lancioni, C.; Carbone, L.; Okhovat, M.
Show abstract
Wildfire smoke is an escalating global public health threat exposing millions of people, including children, to hazardous air pollution each year. Although wildfire smoke toxicants have been linked to a range of adverse health outcomes, including immune dysregulation, the long-term consequences of real-world pediatric wildfire smoke exposure on health and development remain largely unknown. To investigate the persistent effects of early-life exposure on immune health, here we leveraged a cohort of rhesus macaques that experienced nine consecutive days of hazardous wildfire smoke exposure in infancy during the 2020 Oregon Labor Day wildfires. By integrating ex vivo immune stimulations, multiplex cytokine profiling, single-cell transcriptomics, and genome-wide DNA methylation profiling, we identified persistent immunological consequences across molecular and functional levels. We found that a single severe postnatal exposure, in the first three months of life, was associated with persistent change in the innate immune response, including reduced pro-inflammatory cytokine response to a bacterial endotoxin, with subtle but consistent transcriptional changes in myeloid cells, particularly among males. Wildfire smoke exposure was also associated with changes in proportion of B and T/NK cells, and within the T/NK cell compartment, exposed animals exhibited an expansion of cytotoxic cells. Consistent with this, CD8+ T cells displayed extensive transcriptional remodeling and shifted toward more differentiated effector states, with the greatest differentiation observed in animals exposed at the youngest ages. Genome-wide DNA methylation profiling identified smoke-associated methylation changes consistent with acceleration of epigenetic aging, as well as persistent epigenetic alterations impacting genes involved in oxidative stress responses, innate immunity, T cell differentiation, and hematopoiesis. These findings demonstrate that a single severe wildfire smoke exposure during a critical developmental window is associated with extensive immune and epigenetic remodeling that persist years after exposure, providing new insight into the long-term biological consequences of early-life wildfire smoke exposure.
Qian, Z.; Khera, A.; Makhnoon, S.; Chapman, B. E.; Bryant, B.; Sayers, M.; Compton, F.; Eason, S.; Xing, C.; Ahmad, Z.
Show abstract
Background. Cardiovascular-kidney-metabolic (CKM) syndrome affects nearly 90% of US adults, yet most individuals at early, modifiable stages remain unidentified outside clinical care. Blood donation centers offer a scalable, non-clinical venue for CKM screening, but the potential benefit of screening in this context remains unclear. We projected the population-level impact of effective digital return of results (ROR) to inform the design of a pragmatic trial. Methods. We developed a Monte Carlo simulation (100,000 iterations) of the incident major adverse cardiovascular events (MACE), end-stage renal disease (ESRD), and type 2 diabetes (T2DM) preventable by ROR-prompted, guideline-concordant follow-up among donors in CKM Stages 1-2. The estimand counts only events averted by donors who act because of ROR; the intervention effect was modeled directly on strictly positive support, and action was translated into prevented events through a hazard-based cumulative-incidence difference that counts each donor at most once. We evaluated 18 design cells (donor volumes 300,000, 1 million, and 8 million/year; 5- and 10-year horizons; action-rate gains of +10, +20, and +30 percentage points [pp]) and, in a complementary two-arm simulation, the assurance (expected power) of detecting the effect in a single deployment. Results. Under the primary +20 pp scenario, ROR at a single large blood center (300,000 donors/year) is projected to prevent a median of 2,201 events (95% uncertainty interval [UI], 1,099-4,364) over 10 years, scaling to 58,526 (29,154-116,769) at the national donor pool. All 18 design cells had strictly positive 95% lower bounds. The number needed to screen was 136 and the screening cost $2,045 per event prevented (at $15/donor), both invariant to donor volume. Impact scaled linearly with volume and effect size but sub-linearly with the horizon. Detection of the effect was effectively certain at gains of +20 pp or larger (assurance [≥]99.6% in every cell and >99.9% in all but the smallest 5-year cell). Conclusions. Even under the conservative scenario, digital CKM ROR at blood donation centers is projected to prevent hundreds to tens of thousands of incident cardiometabolic events at a screening cost per event well within accepted prevention benchmarks, providing prospective, quantitative justification for a pragmatic, randomized evaluation of digital ROR in non-clinical screening settings.
Page, S.; Easey, K.; Sedgewick, F.; Rai, D.; Stergiakouli, E.
Show abstract
A body of research suggests that autistic individuals are less likely to drink alcohol than neurotypicals. However, emerging studies support a link between autism and alcohol use. This complex relationship is also reflected in studies that have examined the genetic overlap between the two traits. However, it is unclear whether there is a direct causal relationship between them. To explore this, we applied a combination of polygenic score and Mendelian randomisation analyses using publicly available genome-wide summary statistics and phenotypic measures of autism and alcohol consumption from UK Biobank. LD score regression analyses did not provide evidence of a genetic correlation between genetic liability for autism and drinks consumed per week (rg=-0.08; CI95%=-0.19, 0.03). Further, findings from polygenic score analyses did not support an association between genetic liability for autism and overall monthly alcohol intake. Univariable Mendelian randomisation analyses showed little evidence for a total effect of autism, attention deficit hyperactivity disorder (ADHD) or depression on overall monthly alcohol consumption. Multivariable Mendelian randomisation analyses also showed little evidence of a direct effect of autism on drinks per week when controlling for ADHD and depression. It is plausible that genetic liability for autism does not directly increase the amount of alcohol consumed but instead operates via commonly co-occurring difficulties in the autistic community. However, our findings may be due to methodological shortcomings, including weak instruments biasing effects towards to the null. Consequently, results should be interpreted with caution and further research conducted to address these issues.
Bowness, J. S.; Bernal Martinez, A.; Barinka, J.; Schulte-Schrepping, J.; Renders, S.; Waclawiczek, A.; Leppa, A.-M.; Trumpp, A.; Raffel, S.; Haas, S.; Velten, L.
Show abstract
To sustain blood formation, hematopoietic stem and progenitor cells (HSPCs) coordinate a multitude of cell biological processes, from cell cycle control and stress responses to lineage priming. While many genetic regulators of high-level HSPC function have been identified, how HSPCs coordinate more basal cell biological programs, and how such programs relate to stem cell function, remains incompletely understood. Here we use Perturb-seq to profile the transcriptional consequences of targeting 520 genes by CRISPRi in primary mouse HSPC cultures. We developed an analytical strategy to separate perturbation-induced changes in cell-state abundance and clonal heterogeneity from cell-state-local transcriptional effects. From these local perturbation signatures, we identified 19 gene regulatory programs (GRPs) that are defined by co-regulation in response to genetic perturbation, in contrast to co-expression or human curation, and align well with cell biological processes. By decomposing gene expression data from functional and clinical studies into program activity, we show that GRP activities associate with, and predict, phenotypes such as clonal output after transplantation, as well as survival and drug response in retrospective acute myeloid leukemia (AML) cohorts. Together, our study establishes perturbation-derived co-regulation programs as an interpretable framework for linking genetic regulators, cell-biological processes and stem-cell-associated phenotypes.
Orfano, A.; Cisse, A.; Guo, Y.; Han, L.; Fikadu, N.; Thiam, L. G.; Ba, A.; Li, R.; Pouye, M. N.; Mangou, K.; Moore, A. J.; Sene, S. D.; Diallo, F.; Ngom, E. M.; Sadio, B.; Mbengue, A.; Membi, C.; Ngasala, B.; Bazie, T.; Some, F. A.; Olson, N.; Patel, S. D.; Shapiro, L.; Parikh, S.; Foy, B. D.; Cappello, M.; Vigan-Womas, I.; Premji, Z.; Dabire, R. K.; Ouedraogo, J.-B.; Sheng, Z.; Bei, A. K.
Show abstract
Transmission-blocking vaccines (TBVs) are a promising strategy to reduce malaria transmission by targeting parasite stages within the mosquito. However, parasite genetic diversity may limit vaccine efficacy. We used next-generation amplicon deep sequencing to identify non-synonymous single nucleotide polymorphisms (SNPs) in Pfs25 from 184 Plasmodium falciparum isolates from Senegal, Tanzania, Ghana, and Burkina Faso. Prioritized SNPs were introduced into P. falciparum via CRISPR-Cas9. For the G116C variant, gametocyte development was evaluated by microscopy and qPCR, and mosquito infectivity was assessed by SMFAs. We identified 26 SNPs, including 24 novel variants. Functional assays showed that the Pfs25 G116C mutation did not affect gametocyte development or exflagellation. SMFA showed no significant differences in oocyst prevalence or intensity between mutant and WT parasites. These findings highlight the importance of integrating genetic surveillance with functional validation to guide the development of effective transmission blocking interventions
Tindall, C.; Long, R. A.; Naughton, B.; Mapes, B. M.; Vismer, D.; Skinner, H. G.; Malenfant, J.; Maurya, M. R.; Nalls, M. A.; Ramachandran, S.; Nguyen, T.; Peters, M. A.; Scheuermann, R. H.
Show abstract
SysBio FAIRplex is a Common Fund Venture Program that catalogs and indexes data from the Accelerating Medicines Partnership(R) (AMP(R)) Program through a federated model in which data hosts retain custody of their datasets. The central piece of this work is the SysBio Common Data Model (SysBio CDM). AMP is a precompetitive public-private partnership started in 2014 that unites the resources of NIH and private partners to improve our understanding of disease pathways and transform current models for developing new treatments by: - identifying new targets, biomarkers, and development paradigms; - developing leading-edge tools and technologies; - collecting large-scale datasets and supporting analytics for open analysis by the public; and - generating consensus platforms and procedures. A multidisciplinary Task Force was chartered to design the SysBio CDM by extending the Observational Medical Outcomes Partnership (OMOP) Common Data Model into the -omics domain. The Task Force produced a Minimum Viable Product comprising nine OMOP tables; four extension tables for assay and file metadata; and a Common Data Element (CDE) Registry to specify field semantics. This manuscript describes the deliverable: the underlying design choices, the criteria applied in selecting and constructing the extension tables, how the extended model supports multimodal data integration across AMP projects, and what further work to support additional -omics modalities would entail. As an auxiliary methodology, the paper also describes the AI-assisted CDE harmonization workflow used to populate the model.
Iliadis, I.; Heitland, I.; Hoeper, K.; Witte, T.; Kahl, K. G.; Stapel, B.; Meyer-Olson, D.
Show abstract
Objective: The Brief-cope questionnaire explore coping behavior. However, the underlying factor structure remains a subject of ongoing debate. Exploratory factor analyses (EFA) conducted across different populations have identified factor solutions ranging from two to fourteen factors. As of yet, the underlying factor structure of the Brief-cope has not been investigated in patients with seropositive rheumatoid arthritis (RA). Therefore, the aim of this study was to explore the underlying factor structure of the Brief-cope in a German population of seropositive RA. Methods: 216 outpatients with seropositive RA completed the Brief-cope. An EFA with principal axis factoring and Promax rotation was conducted. Results: EFA indicated a five-factor solution. The five-factor solution explained 51.95% of variance. The identified factors were: (1) problem-focused coping (Cronbach's = .851), (2) emotion-focused coping ( = .754), (3) maladaptive coping ( = .747), (4) religious coping ( = .851), and (5) substance-use coping ( = .869). Conclusion: A five-factor solution provided the most appropriate representation of the underlying factor structure of the Brief-cope in patients with seropositive RA. This factor structure may serve as a suitable basis for future analyses of Brief-cope data in comparable RA populations.
Luo, X.; Syreeni, A.; Hill, C.; Smyth, L. J.; Dahlstrom, E. H.; Mutter, S.; Chen, Z.; Natarajan, R.; Pan, S.; Parton, A.; Jackson, H.; McKay, G.; Susztak, K.; Hirschhorn, J. N.; Florez, J. C.; Maxwell, A. P.; Groop, P.-H.; McKnight, A. J.; Sandholm, N.
Show abstract
Hyperglycaemia is a hallmark of diabetes and a major risk factor for diabetic kidney disease (DKD). However, the molecular consequences of long-term cumulative hyperglycaemia (CH) remain unclear. As a stable epigenetic modification, DNA methylation may capture past glycaemic exposure. Here, we assessed CH-associated DNA methylation in 1,245 participants with type 1 diabetes (T1D) from Finland and the United Kingdom-Republic of Ireland cohorts. We identified 17 CH-associated CpGs, with the strongest association at cg19693031 (TXNIP). Longitudinal analyses demonstrate that these CH-associated DNA methylation levels remain stable despite short-term glycaemic fluctuations, suggesting lasting epigenetic imprints of earlier metabolic control. Integrative analyses combining genomic, epigenetic, and proteomic data characterized these CpGs and potential target proteins. Mendelian randomization suggested a causal association between cg20853880 (KLF11) and DKD, supported by chromatin accessibility and kidney KLF11 expression. Our findings suggest that epigenetic changes contribute to metabolic memory and may mediate the effects of hyperglycaemia on DKD.
Hasan, A.; Demidova, E. V.; Priyadarshini, P.; Czyzewicz, P.; Gathuka, L.; Murayama, T.; Zhou, Y.; Kiss, Z. A.; Shastry, R. K.; Andrake, M.; Hearne, G.; Devarajan, K.; Wu, C.; Shah, A.; Schultz, B. M.; Connolly, D. C.; Rosen, G. L.; Canadas, I.; Liu, J. C.; Burtness, B. A.; Smith, J. J.; Dunbrack, R. L.; Golemis, E. A.; Whetstine, J. R.; Meyer, J. E.; Arora, S.
Show abstract
Chemoradiotherapy (CRT) is the standard-of-care therapy for many solid malignancies, yet predictive biomarkers of treatment response remain limited. We identified a germline single nucleotide polymorphism (SNP) in an intrinsically disordered region of the lysine demethylase KDM3C/JMJD1C (p.S464T) that is associated with CRT outcomes in locally advanced rectal cancers (LARC) and head and neck squamous cell carcinoma (LA-HNSCC). In silico modeling with AlphaFold predicted S464T substitution influenced interaction between phosphorylated KDM3C and RNF8 FHA domain. In cellular models, conversion of S464 to T464 increased sensitivity to DNA-damaging agents. S464T substitution impaired damage-induced MDC1-RAP80 signaling and downstream RAP80-BRCA1 colocalization. SNP carrying cells impaired DNA repair causing genotoxic stress that is associated with increased cGAS-cGAMP innate immune signaling and increased apoptosis. Population analyses with the SNP highlighted an increase incidence of UV-induced skin and other cancers, linking inherited variation in the chromatin regulatory gene KDM3C to genome instability, cancer risk, and therapeutic vulnerability.